NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

FastPDB: Towards Bag-Probabilistic Queries at Interactive Speeds

https://doi.org/10.1145/3709691

Huber, Aaron; Kennedy, Oliver; Rudra, Atri; Zhao, Zhuoyue; Feng, Su; Glavic, Boris (February 2025, Proceedings of the ACM on Management of Data)

Probabilistic databases (PDBs) provide users with a principled way to query data that is incomplete or imprecise. In this work, we study computing expected multiplicities of query results over probabilistic databases under bag semantics which has PTIME data complexity. However, does this imply that bag probabilistic databases are practical? We strive to answer this question from both a theoretical as well as a systems perspective. We employ concepts from fine-grained complexity to demonstrate that exact bag probabilistic query processing is fundamentally less efficient than deterministic bag query evaluation, but that fast approximations are possible by sampling monomials from a circuit representation of a result tuple's lineage. A remaining issue, however, is that constructing such circuits, while in PTIME, can nonetheless have significant overhead. To avoid this cost, we utilize approximate query processing techniques to directly sample monomials without materializing lineage upfront. Our implementation inFastPDBprovides accurate anytime approximation of probabilistic query answers and scales to datasets orders of magnitude larger than competing methods.
more » « less
Full Text Available
Drag, Drop, Merge: A Tool for Streamlining Integration of Longitudinal Survey Instruments

https://doi.org/10.1145/3665939.3665965

Pokharel, Pratik; Lee, Juseung; Kennedy, Oliver; Markatou, Marianthi; Talal, Andrew; Good, Jeff; Mukhopadhyay, Raktim (June 2024, ACM)

We explore data management for longitudinal study survey instruments: (i) Survey instrument evolution presents a unique data integration challenge; and (ii) Longitudinal study data frequently requires repeated, task-specific integration efforts. We present DDM (Drag, Drop, Merge), a user interface for documenting relationships among attributes of source schemas into a form that can streamline subsequent efforts to generate task-specific datasets. DDM employs a "human-in-the-loop" approach, allowing users to validate and refine semantic mappings. Through a simulation of user interactions with DDM, we demonstrate its viability as a way to reduce cognitive overhead for longitudinal study data curators.
more » « less
Full Text Available
Overlay Spreadsheets

https://doi.org/10.1145/3597465.3605220

Kennedy, Oliver; Glavic, Boris; Brachmann, Michael (June 2023, HILDA '23: Proceedings of the Workshop on Human-In-the-Loop Data Analytics)

Full Text Available
Efficient Approximation of Certain and Possible Answers for Ranking and Window Queries over Uncertain Data

https://doi.org/10.14778/3583140.3583151

Feng, Su; Glavic, Boris; Kennedy, Oliver (February 2023, Proceedings of the VLDB Endowment)

Uncertainty arises naturally in many application domains due to, e.g., data entry errors and ambiguity in data cleaning. Prior work in incomplete and probabilistic databases has investigated the semantics and efficient evaluation of ranking and top-k queries over uncertain data. However, most approaches deal with top-k and ranking in isolation and do represent uncertain input data and query results using separate, incompatible data models. We present an efficient approach for under- and over-approximating results of ranking, top-k, and window queries over uncertain data. Our approach integrates well with existing techniques for querying uncertain data, is efficient, and is to the best of our knowledge the first to support windowed aggregation. We design algorithms for physical operators for uncertain sorting and windowed aggregation, and implement them in PostgreSQL. We evaluated our approach on synthetic and real world datasets, demonstrating that it outperforms all competitors, and often produces more accurate results.
more » « less
Full Text Available
Runtime provenance refinement for notebooks

https://doi.org/10.1145/3530800.3534535

Deo, Nachiket; Glavic, Boris; Kennedy, Oliver (June 2022, Proceedings of the 14th International Workshop on the Theory and Practice of Provenance)

Full Text Available
TreeToaster: Towards an IVM-Optimized Compiler

https://doi.org/10.1145/3448016.3459244

Balakrishnan, Darshana; Nuessle, Carl; Kennedy, Oliver; Ziarek, Lukasz (June 2021, SIGMOD '21: International Conference on Management of Data)
null (Ed.)
A compiler’s optimizer operates over abstract syntax trees (ASTs), continuously applying rewrite rules to replace sub- trees of the AST with more efficient ones. Especially on large source repositories, even simply finding opportunities for a rewrite can be expensive, as optimizer traverses the AST naively. In this paper, we leverage the need to repeatedly find rewrites, and explore options for making the search faster through indexing and incremental view maintenance (IVM). Concretely, we consider bolt-on approaches that make use of embedded IVM systems like DBToaster, as well as two new approaches: Label-indexing and TreeToaster, an AST- specialized form of IVM. We integrate these approaches into an existing just-in-time data structure compiler and show experimentally that TreeToaster can significantly improve performance with minimal memory overheads.
more » « less
Full Text Available
Efficient Uncertainty Tracking for Complex Queries with Attribute-level Bounds

https://doi.org/10.1145/3448016.3452791

Feng, Su; Glavic, Boris; Huber, Aaron; Kennedy, Oliver A. (June 2021, SIGMOD '21: International Conference on Management of Data)
null (Ed.)
Incomplete and probabilistic database techniques are principled methods for coping with uncertainty in data. unfortunately, the class of queries that can be answered efficiently over such databases is severely limited, even when advanced approximation techniques are employed. We introduce attribute-annotated uncertain databases (AU-DBs), an uncertain data model that annotates tuples and attribute values with bounds to compactly approximate an incomplete database. AU-DBs are closed under relational algebra with aggregation using an efficient evaluation semantics. Using optimizations that trade accuracy for performance, our approach scales to complex queries and large datasets, and produces accurate results.
more » « less
Full Text Available
Reducing Ambiguity in Json Schema Discovery

https://doi.org/10.1145/3448016.3452801

Spoth, William; Kennedy, Oliver; Lu, Ying; Hammerschmidt, Beda; Liu, Zhen Hua (June 2021, SIGMOD '21: International Conference on Management of Data)
null (Ed.)
Ad-hoc data models like Json simplify schema evolution and enable multiplexing various data sources into a single stream. While useful when writing data, this flexibility makes Json harder to validate and query, forcing such tasks to rely on automated schema discovery techniques. Unfortunately, ambiguity in the schema design space forces existing schema discovery systems to make simplifying, data-independent assumptions about schema structure. When these assumptions are violated, most notably by APIs, the generated schemas are imprecise, creating numerous opportunities for false positives during validation. In this paper, we propose Jxplain, a Json schema discovery algorithm with heuristics that mitigate common forms of ambiguity. Although Jxplain is slightly slower than state of the art schema extractors, we show that it produces significantly more precise schemas.
more » « less
Full Text Available
DataSense: Display-Agnostic Data Documentation

Kumari, Poonam; Brachmann, Michael; Kennedy, Oliver; Feng, Su; Glavic, Boris (January 2021, Conference on Innovative Data Systems Research)
null (Ed.)
Documentation of data is critical for understanding the semantics of data, understanding how data was created, and for raising awareness of data quality problem, errors, and assumptions. However, manually creating, maintaining, and exploring documentation is time consuming and error prone. In this work, we present our vision for display-agnostic data documentation (DAD), a novel data management paradigm that aids users in dealing with documentation for data. We introduce DataSense, a system implementing the DAD paradigm. Specifically, DataSense supports multiple types of documentation from free form text to structured information like provenance and uncertainty annotations, as well as several display formats for documentation. DataSense automatically computes documentation for derived data. A user study we conducted with uncertainty documentation produced by DataSense demonstrates the benefits of documentation management.
more » « less
Full Text Available
DataSense: Display Agnostic Data Documentation

Kumari, Poonam; Brachmann, Michael; Kennedy, Oliver; Feng, Su; Glavic, Boris (January 2021, Conference on Innovative Data Systems Research)
Boncz, Peter; Ozcan, Fatma; Patel, Jignesh (Ed.)
Documentation of data is critical for understanding the semantics of data, understanding how data was created, and for raising aware- ness of data quality problem, errors, and assumptions. However, manually creating, maintaining, and exploring documentation is time consuming and error prone. In this work, we present our vi- sion for display-agnostic data documentation (DAD), a novel data management paradigm that aids users in dealing with documenta- tion for data. We introduce DataSense, a system implementing the DAD paradigm. Specifically, DataSense supports multiple types of documentation from free form text to structured information like provenance and uncertainty annotations, as well as several display formats for documentation. DataSense automatically computes documentation for derived data. A user study we conducted with uncertainty documentation produced by DataSense demonstrates the benefits of documentation management.
more » « less
Full Text Available

« Prev Next »

Search for: All records